DataFrame Indexing
The index is an ordered collection of labels used for selection and alignment. It does not have to be a database primary key, but uniqueness and ordering affect many operations.
Core indexers
import pandas as pd
df = pd.DataFrame({
"order_id": ["order-42", "order-43"],
"status": ["paid", "pending"],
"customer_id": ["C1", "C2"],
"amount": [120, 80],
}, index=["order-42", "order-43"])
df.loc["order-42", ["status", "amount"]] # labels
df.iloc[0:10, [1, 3]] # positions
df.at["order-42", "status"] # one scalar by label
df.iat[0, 1] # one scalar by position
Label slices with .loc are generally inclusive at both endpoints when the
labels are present; positional slices with .iloc follow normal Python
half-open semantics.
Set and reset
indexed = df.set_index("order_id")
if not indexed.index.is_unique:
raise ValueError("order_id must be unique")
plain = indexed.reset_index()
Check the resulting index's is_unique property and reject duplicates when uniqueness is required. The verify_integrity argument to set_index is deprecated in pandas 3.0. Sort an index when later operations depend on ordered slicing or time-series behavior.
MultiIndex
A MultiIndex represents multiple label levels on an axis. It is useful when
hierarchical selection, reshaping, or grouped output is central:
sales = pd.DataFrame({
"country": ["Canada", "Canada"], "city": ["Toronto", "Ottawa"], "amount": [10, 20]
})
by_region = sales.set_index(["country", "city"]).sort_index()
toronto = by_region.loc[("Canada", "Toronto")]
canada = by_region.loc["Canada"]
Prefer ordinary columns when the hierarchy is temporary or when downstream
tools expect flat tables. reset_index returns levels to columns.
Alignment hazard
Assignments and arithmetic involving pandas objects align by label. If position is intended, make that conversion explicit and verify lengths. Silent alignment is one of pandas' strengths, but also one of its most common sources of unexpected missing values.
Selection shapes and duplicate labels
With the unique labels above, df.loc["order-42", ["status", "amount"]] returns a length-2 Series, df.iloc[0:10, [1, 3]] a (2, 2) DataFrame, and both scalar examples return "paid". Use df.loc[["order-42"], ["status", "amount"]] to keep a (1, 2) table. A positional slice clips at the end; an out-of-range scalar position raises IndexError.
An Index is an ordered collection, not a mathematical set: duplicates are allowed. A scalar-looking label selection can therefore return multiple rows. Check df.index.is_unique and df.columns.is_unique before relying on scalar results. reindex differs from loc: it constructs the requested labels and inserts missing values for absent ones, while strict label lookup raises KeyError.